Papers with rejection rate
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for unlearning in large language models often hallucinate, generate abnormal token sequences, or behave inconsistently, raising safety and trust concerns. |
| Approach: | They propose a formal definition of unlearning honesty that preserves both utility and honesty on retained knowledge and ensures effective forgetting while encouraging the model to acknowledge its limitations. |
| Outcome: | The proposed method achieves highest rejection rate and refusal stability on Q A tasks from the forget set, nearly double the second-best method. |
PredictaBoard: Benchmarking LLM Score Predictability (2025.findings-acl)
Copied to clipboard
Lorenzo Pacchiardi, Konstantinos Voudouris, Ben Slater, Fernando Martínez-Plumed, Jose Hernandez-Orallo, Lexin Zhou, Wout Schellaert
| Challenge: | Large Language Models (LLMs) fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks. |
| Approach: | They propose a framework to evaluate the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets. |
| Outcome: | The proposed framework evaluates the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets. |